Value Functions Factorization With Latent State Information Sharing in Decentralized Multi-Agent Policy Gradients

نویسندگان

چکیده

Value function factorization via centralized training and decentralized execution is promising for solving cooperative multi-agent reinforcement tasks. One of the approaches in this area, QMIX, has become state-of-the-art achieved best performance on StarCraft II micromanagement benchmark. However, monotonic-mixing per agent estimates QMIX known to restrict joint action Q-values it can represent, as well insufficient global state information single value estimation, often resulting suboptimality. To end, we present LSF-SAC, a novel framework that features variational inference-based information-sharing mechanism extra assist individual agents factorization. We demonstrate such latent sharing significantly expand power factorization, while fully still be maintained LSF-SAC through soft-actor-critic design. evaluate challenge outperforms several methods challenging collaborative further set extensive ablation studies locating key factors accounting its improvements. believe new insight lead local estimation deep learning algorithms. A demo video code implementation found at https://sites.google.com/view/sacmm.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Counterfactual Multi-Agent Policy Gradients

Many complex reinforcement learning (RL) problems such as the coordination of autonomous vehicles, network packet delivery, and distributed logistics are naturally modelled as cooperative multiagent systems. However, RL methods designed for single agents typically perform poorly on such tasks, mainly due to that the joint action space of the agents grows exponentially with the number of agents....

متن کامل

Decentralized Inventory Sharing with Asymmetric Information

We study the information asymmetry issues in a decentralized inventory sharing system consisting of a manufacturer and two independent retailers, who privately hold demand information, non-cooperatively place their orders, but cooperatively share inventories with each other. We find that while the manufacturer needs retailers’ mean demand and standard deviation for her wholesale price decision,...

متن کامل

Intelligent Multi-Agent Dynamical Systems and Policy Sharing

The effects of policy sharing between agents in a multi-agent dynamical system has not been studied extensively. I simulate a system of agents optimizing the same task using reinforcement learning, to study the effects of different population densities and policy sharing. Rich behavior emerges, dependent on the varied parameters, giving insight towards the dynamics of the system. We demonstrate...

متن کامل

Multi-Agent Reinforcement Learning and Genetic Policy Sharing

The effects of policy sharing between agents in a multi-agent dynamical system has not been studied extensively. I simulate a system of agents optimizing the same task using reinforcement learning, to study the effects of different population densities and policy sharing. I demonstrate that sharing policies decreases the time to reach asymptotic behavior, and results in improved asymptotic beha...

متن کامل

Decentralized Multi-Agent Navigation Planning with Braids

We present a novel planning framework for navigation in dynamic, multi-agent environments with no explicit communication among agents, such as pedestrian scenes. Inspired by the collaborative nature of human navigation, our approach treats the problem as a coordination game, in which players coordinate to avoid each other as they move towards their destinations. We explicitly encode the concept...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: IEEE transactions on emerging topics in computational intelligence

سال: 2023

ISSN: ['2471-285X']

DOI: https://doi.org/10.1109/tetci.2023.3293193